Stochastic Modeling of High-level Structures in Handwritten Word Recognition
نویسندگان
چکیده
Handwritten word recognition is an important topic in pattern recognition. It has many applications in automated document processing such as postal address interpretation, bankcheck reading and form reading. There is evidence from psychological studies that word shape plays a significant role in human visual word recognition. High-level structures in handwriting, such as loops, junctions, turns, and ends, are considered to be highly shape-defining. These structures can be more precisely described by their attributes such as position, orientation, curvature, and size. Algorithms based on skeletal graphs are designed to extract structural features. Viewing handwriting as a sequence of structural features, we choose stochastic finite-state automata (SFSAs) as our modeling tool. We extend SFSAs to model high-level structures and their continuous attributes, and view the popular hidden Markov models (HMMs) as special cases of SFSAs obtained by tying parameters on transitions. Experimental results on these two modeling tools have shown advantages of SFSAs over HMMs. To allow real-time applications of the stochastic word recognizers, we introduce several fast-decoding techniques, including character-level dynamic programming, duration constraint, prefix/suffix sharing, choice pruning, etc. A parallel version of the recognizer is also implemented by splitting large lexicons. The resulting word recognizer is better than or comparable to other recognizers in terms of recognition accuracy and speed. For recognizers building word recognition on character recognition, we propose a performance model to associate word recognition accuracy with character recognition accuracy. The model parameters can be determined by multiple regression on accuracy rates obtained on the training data. This model can be used to predict a recognizer’s performance given a lexicon and promises its applications in dynamic classifier selection and combination.
منابع مشابه
Holistic Farsi handwritten word recognition using gradient features
In this paper we address the issue of recognizing Farsi handwritten words. Two types of gradient features are extracted from a sliding vertical stripe which sweeps across a word image. These are directional and intensity gradient features. The feature vector extracted from each stripe is then coded using the Self Organizing Map (SOM). In this method each word is modeled using the discrete Hidde...
متن کاملMixture of Experts for Persian handwritten word recognition
This paper presents the results of Persian handwritten word recognition based on Mixture of Experts technique. In the basic form of ME the problem space is automatically divided into several subspaces for the experts, and the outputs of experts are combined by a gating network. In our proposed model, we used Mixture of Experts Multi Layered Perceptrons with Momentum term, in the classification ...
متن کاملOff-line Arabic Handwritten Recognition Using a Novel Hybrid HMM-DNN Model
In order to facilitate the entry of data into the computer and its digitalization, automatic recognition of printed texts and manuscripts is one of the considerable aid to many applications. Research on automatic document recognition started decades ago with the recognition of isolated digits and letters, and today, due to advancements in machine learning methods, efforts are being made to iden...
متن کاملStochastic trajectory modeling for recognition of unconstrained handwritten words
In this paper we describe an oo-line handwritten word recognition (hwr) system applied to the identi-cation of literal french check amounts. It consists of three successive levels denoted as character, word and phrase level, each of them being related to the previous ones via conditional probability distributions. Training is done on character samples extracted from amount images which are mode...
متن کاملیک روش دو مرحلهای برای بازشناسی کلمات دستنوشته فارسی به کمک بلوکبندی تطبیقی گرادیان تصویر
This paper presented a two step method for offline handwritten Farsi word recognition. In first step, in order to improve the recognition accuracy and speed, an algorithm proposed for initial eliminating lexicon entries unlikely to match the input image. For lexicon reduction, the words of lexicon are clustered using ISOCLUS and Hierarchal clustering algorithm. Clustering is based on the featur...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2007